Hierarchical Multi-Label Text Categorization with Global Margin Maximization

نویسندگان

  • Xipeng Qiu
  • Wenjun Gao
  • Xuanjing Huang
چکیده

Text categorization is a crucial and wellproven method for organizing the collection of large scale documents. In this paper, we propose a hierarchical multi-class text categorization method with global margin maximization. We not only maximize the margins among leaf categories, but also maximize the margins among their ancestors. Experiments show that the performance of our algorithm is competitive with the recently proposed hierarchical multi-class classification algorithms.

برای دانلود رایگان متن کامل این مقاله و بیش از 32 میلیون مقاله دیگر ابتدا ثبت نام کنید

ثبت نام

اگر عضو سایت هستید لطفا وارد حساب کاربری خود شوید

منابع مشابه

Multi-label Text Categorization with Model Combination based on F1-score Maximization

Text categorization is a fundamental task in natural language processing, and is generally defined as a multi-label categorization problem, where each text document is assigned to one or more categories. We focus on providing good statistical classifiers with a generalization ability for multi-label categorization and present a classifier design method based on model combination and F1-score ma...

متن کامل

Exploiting Associations between Class Labels in Multi-label Classification

Multi-label classification has many applications in the text categorization, biology and medical diagnosis, in which multiple class labels can be assigned to each training instance simultaneously. As it is often the case that there are relationships between the labels, extracting the existing relationships between the labels and taking advantage of them during the training or prediction phases ...

متن کامل

A New Kernel-Based Classification Algorithm for Multi-label Datasets

With the emergence of rich online content, efficient information retrieval systems are required. Application content includes rich text, speech, still images and videos. This content, either stored or queried, can be assigned to many classes or labels at the same time. This calls for the use of multi-label classification techniques. In this paper, a new kernel-basedmulti-label classification al...

متن کامل

Multi-label large margin hierarchical perceptron

This paper looks into classification of documents that have hierarchical labels and are not restricted to a single label. Previous work in hierarchical classification focuses on the hierarchical perceptron (Hieron) algorithm. Hieron only supports single label learning. We investigate applying several standard multi-label learning techniques to Hieron. We then propose an extension of the algorit...

متن کامل

Maximal Margin Labeling for Multi-Topic Text Categorization

In this paper, we address the problem of statistical learning for multitopic text categorization (MTC), whose goal is to choose all relevant topics (a label) from a given set of topics. The proposed algorithm, Maximal Margin Labeling (MML), treats all possible labels as independent classes and learns a multi-class classifier on the induced multi-class categorization problem. To cope with the da...

متن کامل

ذخیره در منابع من


  با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید

عنوان ژورنال:

دوره   شماره 

صفحات  -

تاریخ انتشار 2009